npj Genomic Medicine
○ Springer Science and Business Media LLC
Preprints posted in the last 7 days, ranked by how well they match npj Genomic Medicine's content profile, based on 36 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.
Show abstract
Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.
Lee, K. T.; Egleston, B.; Fetzer, D.; Domchek, S. M.; Fleisher, L.; Wen, K.-Y.; Wagner, L.; Roberts, S.; Howe, S.; Cacioppo, C.; Christiansen, J.; Karpink, K.; Selmani, E.; Mastaglio, E.; Weinberg, M.; Wood, E. M.; Feng, J.; John, S.; Schweickert, K.; Mcleod, B.; Bradbury, A. R.
Show abstract
Background: Many at-risk patients lack access to genetic services due to a genetic counselor (GC) workforce shortage. Little is known about how digital alternatives impact patients with and without cancer who meet criteria for genetic testing. Methods: eREACH2 is a randomized 4-arm non-inferiority trial where pre-test (visit 1) and/or return of results (visit 2) GC counseling was replaced with a patient-centered digital intervention. Arms include: A (GC/GC), B (GC/digital), C (digital/GC) and D (digital/digital). Primary outcomes were non-inferiority in uptake of genetic services and change in genetic knowledge and general anxiety from baseline to post-disclosure of results (T0-T2). Secondary cognitive and affective outcomes were assessed using non-inferiority ANOVAs and equivalency chi-squared tests in intention-to-treat and per-protocol analyses. Findings: 773 participants were recruited nationwide; 46.6% from rural areas. Mean age was 51 years (range 20-87), 13% male, 12% non-white, 29% had less than a college education, and 33% had a personal history of cancer. 584 (76%) patients completed testing (14% had a positive result, 16% had a VUS). In the primary ITT analyses, we met the non-inferiority for uptake of genetic services and anxiety, but results were inconclusive for knowledge. Secondary outcomes were heterogeneous across arms. Arm C demonstrated consistently favorable effects, while Arms B and D showed less favorable outcomes in select domains (e.g. satisfaction and MICRA). Patients who received positive or VUS results via digital disclosure had significantly higher MICRA scores - indicating greater negative response to testing. Interpretation: In this large, randomized trial of patients with and without cancer, the eREACH intervention was effective for pre-test counseling, but inconclusive for digital disclosure of results. Exploratory analyses suggest that digital delivery could be a reasonable alternative for individuals receiving negative results, while those receiving positive or VUS results may derive some short-term psychosocial benefit from GC disclosure.
Hasan, A.; Demidova, E. V.; Priyadarshini, P.; Czyzewicz, P.; Gathuka, L.; Murayama, T.; Zhou, Y.; Kiss, Z. A.; Shastry, R. K.; Andrake, M.; Hearne, G.; Devarajan, K.; Wu, C.; Shah, A.; Schultz, B. M.; Connolly, D. C.; Rosen, G. L.; Canadas, I.; Liu, J. C.; Burtness, B. A.; Smith, J. J.; Dunbrack, R. L.; Golemis, E. A.; Whetstine, J. R.; Meyer, J. E.; Arora, S.
Show abstract
Chemoradiotherapy (CRT) is the standard-of-care therapy for many solid malignancies, yet predictive biomarkers of treatment response remain limited. We identified a germline single nucleotide polymorphism (SNP) in an intrinsically disordered region of the lysine demethylase KDM3C/JMJD1C (p.S464T) that is associated with CRT outcomes in locally advanced rectal cancers (LARC) and head and neck squamous cell carcinoma (LA-HNSCC). In silico modeling with AlphaFold predicted S464T substitution influenced interaction between phosphorylated KDM3C and RNF8 FHA domain. In cellular models, conversion of S464 to T464 increased sensitivity to DNA-damaging agents. S464T substitution impaired damage-induced MDC1-RAP80 signaling and downstream RAP80-BRCA1 colocalization. SNP carrying cells impaired DNA repair causing genotoxic stress that is associated with increased cGAS-cGAMP innate immune signaling and increased apoptosis. Population analyses with the SNP highlighted an increase incidence of UV-induced skin and other cancers, linking inherited variation in the chromatin regulatory gene KDM3C to genome instability, cancer risk, and therapeutic vulnerability.
Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.
Show abstract
Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.
Sah, B. K.; Li, C.; Li, J.; Zhu, Z.
Show abstract
Background Conversion surgery for stage IV gastric cancer is supported by a pooled overall survival hazard ratio of 0.36 (95% confidence interval 0.32-0.40) and, in the largest international cohort, median survival of 36.7 versus 12.5-13.8 months on chemotherapy. Survival is measured from diagnosis; the median diagnosis-to-gastrectomy interval is 124 days, which patients must survive to be counted surgical. Methods We simulated cohorts of 3,177 stage IV gastric cancer patients from published parameters: background median survival 14.5 months; median diagnosis-to-surgery interval 124 days (category-specific 92-174 days). Surgery had no effect (true hazard ratio 1.00 by construction). Data were analysed as the literature analyses them (exposure fixed at baseline, follow-up from diagnosis), and by time-varying Cox and landmark analysis. Confounding by indication was added in a second scenario. Results Under immortal time bias alone the naive analysis returned a hazard ratio of 0.794 (95% simulation interval 0.743-0.851), median survival 16.8 versus 12.8 months. Time-varying Cox recovered 1.000 and landmark analysis 1.000-1.004. Bias scaled with the interval: 0.849 at 92 days, 0.715 at 174 days. Adding confounding, the naive estimate fell to 0.601 (0.560-0.644) at strength 0.5 and 0.356 (0.323-0.385) at strength 1.5, overlapping the published estimate; median survival 21.9 versus 8.7 months. Correcting immortal time alone left residual bias (hazard ratio 0.439). Conclusions The reported survival advantage of conversion surgery is reproducible where the operation does nothing; published estimates cannot distinguish benefit from bias. Resolving this requires individual patient data analysed with methods that assign person-time correctly, or completion of JCOG2301.
Venkatesh, R.; Deo, R.; Cappola, T.; Penn Medicine BioBank, ; Ritchie, M. D.; Kim, D.
Show abstract
Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia and a major cause of cardioembolic stroke. Although polygenic risk scores (PRS) are well characterized to quantify inherited susceptibility for AF, they provide limited insight into the pathways and tissues underlying genetic risk, which are critical to uncover for individual risk prediction. In this study, we develop a pathway-level multi-omics representation learning framework that converts individual genetic profiles into interpretable biological features by integrating GWAS-derived pathway burden scores with tissue-specific transcriptomic pathway signals. We constructed machine learning models to assess population-level AF risk prediction performance across genomic and transcriptomic tissue contexts; the pathway-based global attention models substantially improved risk prediction performance over PRS and other baselines (AUROC improved from 0.601 to 0.738). Transformer and graph neural network frameworks then assessed individual-level pathway interpretability, revealing heterogeneous contributions from electrical signaling, cardiac development, and DNA repair pathways to AF risk. This added interpretability highlights the potential of this pathway approach to enable more mechanistically informed risk stratification than static PRS by capturing underlying heterogeneity. To independently assess whether prioritized pathways reflected cardiac regulatory biology, we compared pathway rankings with transcriptional effects predicted by the AlphaGenome foundation model. Variants in highly ranked pathways showed significantly greater predicted effects on expression in atrial and ventricular tissues (FDR = 0.032) relative to controls, providing orthogonal evidence that the model identifies biologically relevant mechanisms. Overall, this work reframes polygenic risk from a single measure of susceptibility to tissue-informed pathway mechanisms, providing a framework for interpretable genomic stratification in complex diseases.
Chen, Y.; Puckett, H.; Clarot, G.; Hawkins, B.; Sharp, K.; Todd, D. A.; Lopez, A.; Bertollo, J. R.; Behar, H. E.; Zeithamova, D.; Xie, H.; Verbalis, A.; VanMeter, A. S.; Gaillard, W. D.; Kenworthy, L.; Vaidya, C. J.
Show abstract
Generalization is a key cognitive process that allows humans to flexibly apply prior knowledge to guide new behaviors. Difficulties with generalization and flexibility are observed across neurodevelopmental disorders, especially autism, limiting adaptive function and quality of life. Cognitive-behavioral treatment benefits some but not all autistic individuals. As treatment requires application of learned skills to everyday life, variability in generalization ability may limit intervention success in autism. While cognitive substrates of learning and generalization are well established, their potential for explaining clinical outcomes is not known. Here, we combined a category learning task with computational modelling to distinguish two learning strategies underlying generalization -- prototype abstraction vs. exemplar memorization -- and tested whether individual differences in these learning strategies predicted real-world intervention outcomes in autistic youth. Fifty-four participants completed the category learning task at two pre-intervention timepoints, and then completed Unstuck and On Target:14-22 intervention targeting flexible problem solving, goal setting, and planning. We found that participants who consistently relied on prototype abstraction (N=26) were subsequently more likely to benefit from the intervention, showing improvement in parent- and self-reported flexibility. These findings identify prototype abstraction as a clinically relevant cognitive capacity that may help explain individual differences in intervention response and support the tailoring of interventions. More broadly, they demonstrate the value of linking basic cognitive mechanisms to clinical outcomes and may inform strategies to enhance the effectiveness of cognitive-behavioral interventions for youth with developmental disabilities.
Qian, Z.; Khera, A.; Makhnoon, S.; Chapman, B. E.; Bryant, B.; Sayers, M.; Compton, F.; Eason, S.; Xing, C.; Ahmad, Z.
Show abstract
Background. Cardiovascular-kidney-metabolic (CKM) syndrome affects nearly 90% of US adults, yet most individuals at early, modifiable stages remain unidentified outside clinical care. Blood donation centers offer a scalable, non-clinical venue for CKM screening, but the potential benefit of screening in this context remains unclear. We projected the population-level impact of effective digital return of results (ROR) to inform the design of a pragmatic trial. Methods. We developed a Monte Carlo simulation (100,000 iterations) of the incident major adverse cardiovascular events (MACE), end-stage renal disease (ESRD), and type 2 diabetes (T2DM) preventable by ROR-prompted, guideline-concordant follow-up among donors in CKM Stages 1-2. The estimand counts only events averted by donors who act because of ROR; the intervention effect was modeled directly on strictly positive support, and action was translated into prevented events through a hazard-based cumulative-incidence difference that counts each donor at most once. We evaluated 18 design cells (donor volumes 300,000, 1 million, and 8 million/year; 5- and 10-year horizons; action-rate gains of +10, +20, and +30 percentage points [pp]) and, in a complementary two-arm simulation, the assurance (expected power) of detecting the effect in a single deployment. Results. Under the primary +20 pp scenario, ROR at a single large blood center (300,000 donors/year) is projected to prevent a median of 2,201 events (95% uncertainty interval [UI], 1,099-4,364) over 10 years, scaling to 58,526 (29,154-116,769) at the national donor pool. All 18 design cells had strictly positive 95% lower bounds. The number needed to screen was 136 and the screening cost $2,045 per event prevented (at $15/donor), both invariant to donor volume. Impact scaled linearly with volume and effect size but sub-linearly with the horizon. Detection of the effect was effectively certain at gains of +20 pp or larger (assurance [≥]99.6% in every cell and >99.9% in all but the smallest 5-year cell). Conclusions. Even under the conservative scenario, digital CKM ROR at blood donation centers is projected to prevent hundreds to tens of thousands of incident cardiometabolic events at a screening cost per event well within accepted prevention benchmarks, providing prospective, quantitative justification for a pragmatic, randomized evaluation of digital ROR in non-clinical screening settings.
Yang, Y.; Vasudevaraja, V.; Serrano, J.; Mohamed, H.; Kelly, S.; Jour, G.; Gindin, T.; Park, K.; Jones, D.; Feng, X.; Pinnell, J.; Mclennan, S.; Tin, M. Y.; Tsirigos, A.; Snuderl, M.; Wrzeszczynski, K. O.
Show abstract
Next-generation sequencing (NGS) for the detection of somatic variants has become the method of choice in a variety of molecular oncology fields and in the clinic. Its use ranges from sequencing entire tumor genomes and transcriptomes to targeted clinical diagnostic gene panels. The NYU Langone Genome PACT (Profiling of Actionable Cancer Targets, LG-PACT) assay is a qualitative in vitro diagnostic test that uses targeted next generation sequencing (NGS) of formalin-fixed paraffin-embedded (FFPE) tumor tissue matched with normal specimens from patients to detect gene alterations in a targeted panel covering 606 genes and the TERT promoter. Indications for testing are cancer (solid tumors and hematological malignancies) where a mutational profile from multiple genes would be informative for disease stratification, prognosis, or treatment options including targeted therapies and eligibility for clinical trials. The test is intended to provide information on somatic mutations including point mutations, small insertions/deletions (indels), and copy number aberrations for diagnostic and treatment decisions. LG-PACT is a United States Food and Drug Administration (FDA) cleared diagnostic test (510K: K202304). The clinical interpretation of sequencing data of molecular tumor markers from NGS encompasses automated variant calling tools with human interpretation. This final mostly manual review of data step is intensive, involving highly trained scientists, encompassing literature review, interpretation and clinical tier classification by pathologists, who then provide a complete molecular diagnostic report to the treating oncologists. We provide analysis of 1339 clinical genomic profiles from 31 different cancers and their subtypes, comprising of central nervous system (CNS) 792 (59%) cases (incl. meningioma, glioma and glioblastoma), with 267 (20%) cases predominantly of lung, pancreatic and colorectal and 280 of others (21%). Here, we present the technical challenges of validating an NGS oncological diagnostic targeted assay for clinical grade accuracy and sensitivity for patient care. We show how copy number alterations provide a more comprehensive description of the tumors genomic profile. We then outline the utility of targeted panel sequencing based on certified pathologist selection of reportable variants for our current patient cohort. Where analysis of variant detection has led to 49.4% (661/1339) of our clinical tumor samples containing mutations in known therapy targeted genes, 35.6% (477/1339) with mutation detected in other genes, and 15% (201/1339) cases being negative.
Gao, Y.; Yu, S.; Xia, Y.; Chen, S.; Xia, S.; An, R.; Zeng, J.; Zhao, F.; Ma, Y.; Wang, Y.; Xie, X.; Zhang, J.
Show abstract
Prognostic models in oncology are developed one cancer at a time, from that cancer's own labelled outcomes, and fail where prognostic information is scarcest. Rare cancers account for roughly a fifth of diagnoses and most paediatric malignancies, yet seldom supply enough events for a reliable time-to-event model. We therefore asked whether a representation learned without outcome labels can supply what those cohorts cannot. A Transformer encoder was pretrained by masked field-value modelling on 9425135 tumour records from the SEER 17 registries, diagnosed in 2000 to 2023. Only diagnosis-time fields passing a fail-closed coding-verification gate were admitted, and each record was emitted as an era-specific and a harmonised view, keeping two decades of recoding auditable. The encoder was then frozen and read by a linear Cox head for overall survival. Nine rare cancers were removed from the pretraining corpus entirely, each requiring an independent pretraining run. On a sealed test partition, all nine exceeded an architecture-identical random frozen encoder in Harrell concordance by +0.0034 to +0.0368, every lower confidence limit above zero. At 256 labelled patients, all 67 cancers favoured the pretrained representation over budget-matched Cox regression, median difference +0.0283. The advantage was bounded: given the entire training set, Cox regression was favoured in seven of nine rare cancers. The encoder did not outperform a field-frequency baseline on its own objective, so upstream reconstruction did not predict downstream transfer. Outcome-agnostic registry pretraining carries prognostic signal into cancers it has never seen, and is most useful where labels are fewest, without establishing clinical utility.
Jaholkowski, P.; Parker, N.; Sveen, I. O.; Wistrom, E. D.; Fominykh, V.; Szabo, A.; Parekh, P.; Frei, O.; Smeland, O. B.; O'Connell, K. S.; Djurovic, S.; Dale, A. M.; Shadrin, A. A.; Andreassen, O. A.
Show abstract
Recent large-scale studies have enabled new knowledge about genetic underpinnings of morphological and electrophysiological alterations of the retina. Variation in retinal traits, often of neurodevelopmental origin, have been linked to major psychiatric disorders (MPDs). Here, we investigate the genetic overlap between MPDs and key retinal traits to identify underlying molecular mechanisms. We obtained genome-wide associations studies data for bipolar disorder (BD), major depression (MD), schizophrenia (SCZ), and the retinal traits retinal nerve fibre layer thickness (RNFL), ganglion cell inner plexiform layer thickness (GCIPL), and vertical cup-disc ratio (VCDR). We estimated the number of trait-influencing variants shared between traits with MiXeR and identified shared genetic loci with condFDR. Subsequently, we examined the biological pathways of the genes mapped to shared loci. This revealed that GCIPL shared the most genetic variants with MPDs (~60%), followed by RNFL (~40%), and VCDR (~20%). The genetic variants shared between retinal traits and MPDs showed disorder-specific patterns with more pronounced overlaps of SCZ and BD with RNFL, and MD negatively correlated with GCIPL. Gene-pathway analysis highlighted the importance of GABAergic neurotransmission and a two-stage neurodevelopmental process in SCZ, whereas the role of mitochondria and a weaker developmental component were observed in BD. The results also implicated synaptic functioning and gene-expression processes in MD. Furthermore, polygenic analysis suggested that the genetic architecture of retinal traits can distinguish between MPDs. Our findings indicate shared genetic underpinnings between retinal traits and SCZ, BD, and MD, implicating altered neurodevelopment and neurotransmission underlying the retinal link to major psychiatric disorders.
Page, S.; Easey, K.; Sedgewick, F.; Rai, D.; Stergiakouli, E.
Show abstract
A body of research suggests that autistic individuals are less likely to drink alcohol than neurotypicals. However, emerging studies support a link between autism and alcohol use. This complex relationship is also reflected in studies that have examined the genetic overlap between the two traits. However, it is unclear whether there is a direct causal relationship between them. To explore this, we applied a combination of polygenic score and Mendelian randomisation analyses using publicly available genome-wide summary statistics and phenotypic measures of autism and alcohol consumption from UK Biobank. LD score regression analyses did not provide evidence of a genetic correlation between genetic liability for autism and drinks consumed per week (rg=-0.08; CI95%=-0.19, 0.03). Further, findings from polygenic score analyses did not support an association between genetic liability for autism and overall monthly alcohol intake. Univariable Mendelian randomisation analyses showed little evidence for a total effect of autism, attention deficit hyperactivity disorder (ADHD) or depression on overall monthly alcohol consumption. Multivariable Mendelian randomisation analyses also showed little evidence of a direct effect of autism on drinks per week when controlling for ADHD and depression. It is plausible that genetic liability for autism does not directly increase the amount of alcohol consumed but instead operates via commonly co-occurring difficulties in the autistic community. However, our findings may be due to methodological shortcomings, including weak instruments biasing effects towards to the null. Consequently, results should be interpreted with caution and further research conducted to address these issues.
Ebneabbasi, A.; Warrier, V.; Montagnese, M.; Romero Garcia, R.; Bethlehem, R. A. I.; Rittman, T.
Show abstract
Neighbourhood deprivation is one of the few potential policy-modifiable risk factors for psychiatric and neurological disorders, but the neurobiological pathways underlying these associations remain unclear. We investigated these relationships across three cohorts spanning the life span: the Healthy Brain and Child Development (HBCD) Study (n = 84, aged 0 to 4 weeks postnatal), the Adolescent Brain Cognitive Development (ABCD) Study (n = 4,792, aged 9 to 10 years), and the UK Biobank (UKB; approximately 500,000 adults, aged 44 to 87 years). Neighbourhood deprivation was associated with elevated disease risk, and individual lifestyle factors accounted for only a small fraction of this burden, indicating that the much larger residual effect reflects broader contextual characteristics of deprived environments rather than individual behaviours alone. Across all cohorts, greater deprivation consistently predicted lower cortical and subcortical brain volume, with effects detectable in early development and substantially stronger in adulthood. Across disorders, regional brain volume emerged as a consistent neuroanatomical mediator linking neighbourhood deprivation to neuropsychiatric disease. We further showed that deprivation preferentially affects brain regions intrinsically vulnerable to neuropsychiatric disorders. Spatial decoding analyses implicated dopaminergic and serotonergic neurotransmitter systems together with specific excitatory and inhibitory neuronal classes. Importantly, both the deprivation effects and their neuroanatomical mediation patterns were replicated across independent populations. Our study delivers a translational framework linking neighbourhood deprivation to brain health, which could inform public health policies and preventive interventions.
Zhao, L.; Zeng, Y.; Abelman, D. D.; Lin, W.; Luo, P.
Show abstract
Motivation: Cell-free DNA methylation provides a minimally invasive signal for early cancer detection and tissue-of-origin prediction. Most methods represent methylation measurements as independent fixed-window features and therefore do not explicitly model relationships among genomic regions. Results: We developed PANGEM (Pan-cancer Graph-based Cancer Detection Using the Cell-free DNA Methylome), a graph-learning framework that represents genomic bins as nodes and integrates CpG context, genomic proximity, and sample-specific methylation similarity in the graph topology. Across five repeated stratified train-test splits, PANGEM achieved the highest mean performance among evaluated methods, with an AUROC/AUPR of 0.997/1.000 for binary cancer detection and macro-AUROC/AUPR of 0.977/0.870 for multiclass tissue-of-origin prediction. In the independent INSPIRE cohort, 72 of 78 cancer cases (92.3%) exceeded the binary classification threshold, and PANGEM correctly classified 9 of 17 head and neck cancer cases (52.9%), the highest accuracy among evaluated methods. Subnetwork analysis further identified recurrent, graph-connected methylation patterns, including a 111-DMR subnetwork with increased methylation in cancer samples.
Alquicira-Hernandez, J.; Dorans, E.; Tomofuji, Y.; Nathan, A.; Raychaudhuri, S.
Show abstract
Single-cell technologies enable linking disease-risk variants to gene regulatory effects in specific cell-state contexts. However, most so called "single-cell eQTL" studies use a "pseudobulking" strategy to identify expression Quantitative Trait Loci (eQTLs), obscuring subtle dynamic regulatory effects of disease alleles. Here, we propose Dynema (Dynamic eQTL mapping in single cells) for fast and accurate genome-wide mapping of context-dependent and independent eQTL effects at true single-cell resolution. To identify eQTLs, Dynema uses a Poisson model with cluster robust variance estimators (CRVEs) to account for correlation of single-cell profiles from the same individual. In contrast to other common methods, Dynema achieves statistical calibration and scales to genome-wide analysis in large single-cell datasets in realistic timeframes. We applied Dynema to two independent T cell datasets and identified reproducible cell-state-dependent eQTL effects. Some cell-state-dependent eQTLs are missed by pseudobulking approaches, and many others are conditionally independent from lead eQTL effects. We show that TSPAN32 and other autoimmune loci colocalize with cell-state-dependent eQTLs. Mapping context-dependent eQTLs at single-cell resolution enables the definition of the molecular effects of complex disease alleles.
Majumder, B. P.; Linak, J. A.; Adamson, R.; Aguilera, R. L.; Agarwal, D.; Reitz, Z.; Loiselle, S.; Devarakonda, S.; Clark, P.; Paulson, K. G.; Stanton, S.
Show abstract
In large data sets discovery is often limited to pre-conceived hypotheses and data fishing. Here we tested whether systematic exploration of AI generated hypotheses could uncover clinically meaningful signals in extensively studied data. We deployed AutoDiscovery, a newly launched large language model (LLM) framework designed to search for hypotheses based on surprisal and systematically interrogate complex datasets, on The Cancer Genome Atlas breast cancer cohort. The system did not identify clinically meaningful novel findings without human input. However, a seeded warm-start run with minimal text input from an oncologist revealed multiple interesting and surprising hypotheses. Among these was that a robust immune signature was present across all subtypes of invasive lobular carcinoma (ILC) that exceeded invasive ductal carcinoma (IDC). This observation was independently validated in independent cohorts and confirmed by high-sensitivity multi-immunofluorescence tumor tissue analyses. These results suggest immunotherapy approaches should be tested in ILC including early stage ER+HER2- ILC; these patients are currently excluded from large neoadjuvant immunotherapy trials. They further demonstrate that surprisal-based hypothesis generation frameworks can extract previously unappreciated patterns from deeply interrogated cancer datasets and imply that disease domain experts working with LLMs can derive more meaningful insights from complex data than either could achieve alone.
Boden-Albala, B.; Wing, J.; Landry, M. J.; Castro, M.; Gutierrez, D.; Cardenas, C.; Rousseau, J.; Rahmani, A. M.; Chavez, A.; Ding, X.; Kurzman, A.; Albala, B.
Show abstract
Background: Cardiovascular disease (CVD) disproportionately burdens underserved communities, where social determinants of health (SDOH) perpetuate persistent disparities. Family-based interventions leveraging social support represent a promising yet understudied approach. We describe the rationale, design, and methods of the Skills-based Educational strategies for the Reduction of Vascular Events in Orange County (SERVE OC) RCT and present baseline characteristics of enrolled families. Methods: SERVE OC is a 2-arm RCT of 190 Latino and Vietnamese families (486 individuals) randomized to the family-based intervention or individual self-management. The intervention was grounded in social network theory while employing community engaged strategies. Primary outcomes include achieving ideal cardiovascular health (CVH) defined by AHA Life's Essential 8 (LE8) and systolic blood pressure reduction at 12, 24, and 36 months. Baseline assessments include demographics, LE8, psychosocial factors, food security, and SDOH. Descriptive statistics and regression analyses examined cohort characteristics and associations between SDOH, food security, and LE8. Results: Over 83% of participants had suboptimal LE8 scores. Average adult total LE8 scores were 66.61 {plus minus}11.96, with physical activity as the weakest domain, compared to an average of 76.52{plus minus}10.15 in children. Greater SDOH burden and food security were associated with significantly lower odds of ideal CVH and lower LE8 scores respectively. Conclusions: SERVE OC demonstrates the feasibility of enrolling families in community-engaged RCT targeting CVD disparities in underserved population. Baseline findings confirm substantial CVD risk and SDOH burden underscoring the need for multi-level, culturally tailored interventions. Trials results will inform scalable, family-focused strategies for CVD prevention across the life course. Clinical Trial Registration: URL: https://www.clinicaltrials.gov/; Unique Identifier: NCT05641519.
Kouam, C.; Mingle, J.; Alvarez Jerez, P.; Evans, A.; Moller, A.; Baker, B.; Weller, C.; Paquette, K.; Brooks, J.; Grant, S. M.; Ayuketah, A.; Meredith, M.; Palade, J.; Malik, L.; Hise, K.; Raphael Gibbs, J.; Anderson, J.; Ding, J.; Harbert, R.; Fu, Y.; Zheng, X.; Garcia-Ruiz, S.; Gustavsson, E. K.; Blauwendraat, C.; Ryten, M.; Sedlazeck, F.; Ferrucci, L.; Reed, X.; Nalls, M. A.; Cookson, M. R.; Van Keuren-Jensen, K.; Hutchins, E.; Jain, M.; Billingsley, K. J.
Show abstract
Isoform-resolved transcriptomics is fundamental to decoding the molecular complexity of the human brain, yet population-scale long-read RNA sequencing has remained inaccessible due to labor-intensive library preparation, sensitivity to RNA degradation in postmortem tissue, and the absence of integrated, reproducible analysis pipelines. Here we present SALRR (Scalable Analysis of Long-Read RNA-seq), an integrated wet-lab and computational platform designed to overcome these barriers. Automated ONT long-read cDNA library preparation on the Hamilton Microlab NGS STAR platform reduces hands-on time by 67% and enables 24 libraries per operator per day while maintaining performance across RNA integrity values. A modular, Snakemake-based pipeline performs end-to-end processing from ONT signal data to isoform-level quantification, incorporating SIRV spike-in calibration, multi-stage quality control, and stringent isoform validation. Applied to 10 postmortem frontal cortex samples from the North American Brain Expression Consortium, SALRR identified 31,607 high-confidence isoforms from 10,075 genes, including 8,532 novel splice variants absent from GENCODE v49, and complex splicing events systematically missed by short-read sequencing at neurodegeneration-relevant loci, including GBA1, CCNF, CHCHD10, and TREM2. All protocols and code are openly available, providing a scalable, community-ready framework for isoform-resolved transcriptomics in neurodegeneration, aging, and complex brain disease.
Efthymiou, S.; Tabata, K.; Dafsari, H. S.; Schober, E.; Latza, C.; Isaoglu, M.; Abuelrub, A.; Rad, A.; Firoozfar, Z.; Turchetti, V.; Lin, R. Q.; Maroofian, R.; Wiethoff, S.; Afzal, E.; Zafar, F.; Rana, N.; McRae, A. M.; Kaiyrzhanov, R.; Guliyeva, U.; Gulieva, S.; Melikishvili, G.; Lespinasse, J.; Vitobello, A.; Denomme-Pichon, A.-S.; Wentzensen, I. M.; Mefford, H. C.; Briere, L. C.; A Walker, M.; A High, F.; Sweetser, D. A.; Kendall, M.; Franchi, M.; Brown, M.; Latner, D.; Joset, P.; Ivanovski, I.; Alfadhel, M.; Alluhaydan, I.; Frederiksen, A. S.; Arriens, V.; Hanker, B.; Mankad, K.; Guerin, J
Show abstract
Pathogenic variants in RUBCN, encoding the Run domain Beclin-1 interacting and cysteine-rich domain-containing protein (Rubicon) have been implicated in autosomal recessive spinocerebellar ataxia 15 (SCAR15). However, the molecular mechanisms underlying disease pathogenesis remain poorly understood. Here, we report 18 individuals from 15 unrelated families harbouring biallelic RUBCN variants, who present with an aggressive neurodevelopmental disorder variably characterized by seizures, developmental delay, intellectual disability and movement abnormalities that cause regression, progressive brain atrophy and neurodegenerative features. Through functional characterization, we demonstrate that a subset of disease-associated putative truncating variants disrupt autophagy regulation. In Caenorhabditis elegans models, loss-of-function RUBCN variants result in an increased autophagic flux and impaired neuronal function, recapitulating key features in humans. Correspondingly, cellular assays reveal that nonsense and frameshift RUBCN variants lead to defective autophagy inhibition, underscoring a crucial role for RUBCN as a key negative autophagy regulator. Molecular dynamics simulations rank the eleven missense variants by structural effect, with p.Arg813Trp alone altering the target protein at both the local and the regional level and lying within the RAB7A-binding module that the truncating alleles remove altogether. Our findings establish and expand the RUBCN-related disorders as a clinically and molecularly distinct subset of autophagy-related diseases. By delineating both the genetic landscape and cellular consequences of Rubicon dysfunction, this study enhances our understanding of autophagy-related neurodevelopmental disorders and provides a foundation for future therapeutic investigations.
SULAIMAN, M. A.; Oyeyemi, B. F.
Show abstract
Sub-Saharan African populations carry pharmacogenomic alleles poorly represented in the European-derived reference panels underlying most clinical genotyping tools. We present a curated, machine-readable catalog of nine actionable alleles across six pharmacogenes (CYP2D6, CYP2B6, CYP2C9, CYP2C19, CYP3A5, NAT2) with African-specific frequency ranges, functional annotations, and evidence levels derived from reanalysis of 661 high-coverage whole-genome sequences across seven 1000 Genomes Project African populations. Direct comparison against PharmCAT v3.4.0 shows that CYP2D6 produces zero diplotype calls (0/661 samples callable) due to monomorphic reference positions absent from standard variant-only VCF output, a known limitation whose consequences for African allele carriers had not been reported. afripharmagen's reduced-position strategy identifies 243 CYP2D617 and 134 CYP2D629 carriers from the same input. For CYP2B6, CYP2C9, CYP2C19, and NAT2, both tools show concordance of 95-100%. Frequency gradients (CYP2B66: 30-50%; CYP2D617: 15-35% in West Africa; CYP3A5*1: 60-95%) translate directly into prescribing risk for efavirenz, tramadol, tacrolimus, and isoniazid. Pharmacogenomic decision support in African settings must incorporate population-specific allele definitions and input-format-aware strategies.